Papers with alignment paradigms
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration (2024.emnlp-main)
Copied to clipboard
Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher, Chan Young Park, Yejin Choi, Yulia Tsvetkov
| Challenge: | Existing alignment paradigms for large language models learn an averaged human preference and struggle to model diverse preferences across cultures, demographics, and communities. |
| Approach: | They propose a modular framework that "plugs" into a base LLM a pool of smaller but specialized community LMs where models collaborate in distinct modes to support three modes of pluralism: Overton, steerable, and distributional. |
| Outcome: | The proposed framework “plugs into” a base LLM a pool of smaller but specialized community LMs, where models collaborate in distinct modes to support three modes of pluralism: Overton, steerable, and distributional. |
StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work focused on typographic and pixel-level perturbations, leaving the study of SCO unexplored. |
| Approach: | They propose a framework that exploits MLLMs' diagrammatic reasoning capabilities to bypass safety guardrails. |
| Outcome: | The proposed framework exploits the model's reasoning capabilities to bypass safety guardrails. |
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to align Large Language Models with human values model an averaged or monolithic preference, despite progress in pluralistic alignment, no prior work has focused on health . |
| Approach: | They propose a benchmark dataset to assess and benchmark pluralistic alignment methodologies. |
| Outcome: | The proposed model can model pluralistic views within health domains. |
UniCreative: Unifying Long-form Logic and Short-form Sparkle via Reference-Free Reinforcement Learning (2026.findings-acl)
Copied to clipboard
Xiaolong Wei, Zerun Zhu, Simin Niu, Xingyu Zhang, Peiying Yu, Changxuan Xiao, Yuchen Li, Jicheng Yang, Zhejun Zhao, Chong Meng, Long Xia, Daiting Shi
| Challenge: | Existing alignment paradigms for creative writing use static reward signals and supervised data. |
| Approach: | They propose a constraint-aware reward model that synthesizes query-specific criteria to provide fine-grained preference judgments. |
| Outcome: | The proposed framework aligns models with human preferences across content quality and structural paradigms without supervised fine-tuning and ground-truth references. |
QuantumQA: Enhancing Scientific Reasoning via Physics-Consistent Dataset and Verification-Aware Reinforcement Learning (2026.acl-long)
Copied to clipboard
Songxin Qu, Tai-Ping Sun, Yun-Jie Wang, Huan-Yu Liu, Cheng Xue, Xiao-Fan Xu, Han Fang, Yang Yang, Yu-Chun Wu, Guo-Ping Guo, Zhao-Yun Chen
| Challenge: | Large language models lack reliability in scientific domains that require strict adherence to physical constraints. |
| Approach: | They propose a large-scale dataset constructed via a task-adaptive strategy and a hybrid verification protocol that combines deterministic solvers with semantic auditing to guarantee scientific rigor. |
| Outcome: | The proposed model outperforms baselines and general-purpose preference models and is competitive with proprietary models. |